Papers with machine translation training
Towards the First NLP Benchmark for Ladin - an Extremely Low-Resource Language (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are limited in low-resource languages due to lack of labeled training data. |
| Approach: | They propose to use Ladin as a model for sentiment analysis and question answering by incorporating Italian data into machine translation training. |
| Outcome: | The proposed method improves on existing Italian–Ladin translation baselines. |
Machine Translation Models are Zero-Shot Detectors of Translation Direction (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to detect the translation direction of parallel text are lacking in the machine translation community. |
| Approach: | They propose an unsupervised approach to detection of translation direction of parallel texts . they use a simple hypothesis that p(translation|original)>p(original|translation) they confirm the approach is effective for high-resource language pairs . |
| Outcome: | The proposed approach achieves document-level accuracies of 82–96% for NMT-produced translations and 60–81% for human translations, based on the model used. |
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)
Copied to clipboard
Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, Jörg Tiedemann
| Challenge: | a new massive multilingual dataset is available for language modeling and machine translation training. |
| Approach: | They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora . |
| Outcome: | The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 . |